Papers by Rao Muhammad Anwer

8 papers
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts (2025.findings-acl)

Copied to clipboard

Challenge: TimeTravel is a benchmark of 10,250 expert-verified historical artifact samples spanning 266 distinct cultures across 10 major historical regions.
Approach: They evaluate contemporary AI models on TimeTravel, highlighting their strengths and identifying areas for improvement.
Outcome: The timeTravel benchmark covers 266 cultures and 10 major historical regions and aims to establish AI as reliable partner in preserving cultural heritage.
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches do not emphasize step-wise problem-solving.
Approach: They propose a visual reasoning chain benchmark and a fine-grained reasoning metric that evaluates correctness and logical coherence at each step.
Outcome: The proposed framework outperforms existing models in six benchmarks and is 5x faster during inference scaling.
MAviS: A Multimodal Conversational Assistant For Avian Species (2025.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal large language models face challenges when it comes to specialized topics like avian species.
Approach: They propose a large-scale multimodal avian species dataset that integrates image, audio, and text modalities for over 1,000 bird species.
Outcome: The proposed model outperforms the baseline MiniCPM-o-2.6 by a large margin.
BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalities (2025.findings-emnlp)

Copied to clipboard

Challenge: BiMediX2 is a bilingual (Arabic-English) large multimodal model that supports text-based and image-based medical interactions.
Approach: They introduce BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model that supports text-based and image-based medical interactions.
Outcome: The model outperforms existing models by over 9% in English and more than 20% in Arabic evaluations.
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM (2025.findings-acl)

Copied to clipboard

Challenge: Existing speech-enabled LLMs degrade conversational quality by modifying the LLM, compromising its linguistic capabilities.
Approach: They propose a lightweight 30M-parameter, LLM-agnostic, autoregressive streaming TTS system that generates high-quality speech with low latency.
Outcome: The proposed system achieves a significantly lower word error rate compared to speech-enabled LLMs while operating at comparable latency.
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark (2025.findings-naacl)

Copied to clipboard

Challenge: Recent years have witnessed a significant interest in developing large multimodal models capable of performing various visual reasoning and understanding tasks.
Approach: They propose to use Arabic as a language to evaluate large multi-modal models capable of performing visual reasoning and understanding tasks.
Outcome: The proposed benchmark comprises eight diverse domains and 38 sub-domains to represent a large population of over 400 million speakers.
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: a benchmark is designed to assess the comprehension of Arabic poetry by large language models in 12 historical eras.
Approach: They propose a benchmark to assess the comprehension of Arabic poetry by large language models in 12 historical eras.
Outcome: The benchmark assesses the comprehension of Arabic poetry by large language models in 12 historical eras.
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding (2026.eacl-long)

Copied to clipboard

Challenge: a benchmark of 1,272 samples containing about 1,475 unique words is available for Arabic calligraphy . the dataset reflects real-world challenges in Arabic writing, such as calligraphic variation and artistic distortions .
Approach: They evaluated 13 leading Arabic and multilingual multimodal models and paired them with sentence-level annotations to evaluate their calligraphy models.
Outcome: The benchmark evaluates 13 leading Arabic and multilingual multimodal models . it shows they struggle with calligraphic variation, artistic distortions, and precise visual–text alignment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations